Papers with verbatim memorization

2 papers
Copyright Violations and Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: a recent study examines the extent to which language models can memorize training data . a fair use exemption to copyright laws allows for limited use of copyrighted material .
Approach: They examine the extent to which language models can redistribute copyrighted text . they use a range of popular books and coding problems to study copyright violations .
Outcome: This study examines the extent to which language models can redistribute copyrighted text . it shows that language models may memorize entire chunks of training data .
Demystifying Verbatim Memorization in Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that Large Language Models (LLMs) memorize long sequences verbatim, with serious copyright and privacy implications.
Approach: They develop a framework to study verbatim memorization in a controlled setting by continuing pre-training from Pythia checkpoints with injected sequences.
Outcome: The proposed framework creates a control model M () and a treatment model M with injected sequences.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations